Reorder channelwise gated delta rule chunked hot loops for autovectorization (#21021)#21021
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21021
Note: Links to docs will display an error until the docs builds have been completed. ✅ No FailuresAs of commit 6348e75 with merge base 8134bb2 ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D112598714. |
This PR needs a
|
…ization (pytorch#21021) Summary: Reorder the chunked prefill inner loops (steps 1, 4, 5, 6) so the innermost loop runs contiguously over the head dimension (k or v) instead of striding down a column of the state / pv. This lets the compiler autovectorize the now-unit-stride AXPYs; hand-written at::vec was tried and was slower than the compiler output, so the loops stay scalar. Differential Revision: D112598714
b8d62b5 to
dac4505
Compare
…ization (pytorch#21021) Summary: Reorder the chunked prefill inner loops (steps 1, 4, 5, 6) so the innermost loop runs contiguously over the head dimension (k or v) instead of striding down a column of the state / pv. This lets the compiler autovectorize the now-unit-stride AXPYs; hand-written at::vec was tried and was slower than the compiler output, so the loops stay scalar. Reviewed By: billmguo Differential Revision: D112598714
f5857be to
3a6025b
Compare
…ization (pytorch#21021) Summary: Reorder the chunked prefill inner loops (steps 1, 4, 5, 6) so the innermost loop runs contiguously over the head dimension (k or v) instead of striding down a column of the state / pv. This lets the compiler autovectorize the now-unit-stride AXPYs; hand-written at::vec was tried and was slower than the compiler output, so the loops stay scalar. Reviewed By: billmguo Differential Revision: D112598714
…ization (pytorch#21021) Summary: Reorder the chunked prefill inner loops (steps 1, 4, 5, 6) so the innermost loop runs contiguously over the head dimension (k or v) instead of striding down a column of the state / pv. This lets the compiler autovectorize the now-unit-stride AXPYs; hand-written at::vec was tried and was slower than the compiler output, so the loops stay scalar. Reviewed By: billmguo Differential Revision: D112598714
3a6025b to
05aadb3
Compare
…ization (pytorch#21021) Summary: Reorder the chunked prefill inner loops (steps 1, 4, 5, 6) so the innermost loop runs contiguously over the head dimension (k or v) instead of striding down a column of the state / pv. This lets the compiler autovectorize the now-unit-stride AXPYs; hand-written at::vec was tried and was slower than the compiler output, so the loops stay scalar. Reviewed By: billmguo Differential Revision: D112598714
05aadb3 to
d103906
Compare
…BUCK (pytorch#21105) Summary: Relands the two-pass optimization for the `channelwise_gated_delta_rule` custom op (originally pytorch#21020, D112596724), which was reverted in D113048961 because it broke OSS `unittest macos / linux`. The revert was caused by the benchmark BUCK target: ``` runtime.python_binary(name = ..., srcs = [...], main_module = ...) ``` Fix: move the source into a `runtime.python_library` and have the `runtime.python_binary` reference it via `deps` with only `main_module` Differential Revision: D113076546
…h#21061) Summary: Route the channelwise gated delta rule by sequence length: T == 1 keeps the two-pass token recurrence for autoregressive decode, while T != 1 uses a chunkwise WY/UT formulation for prefill. The chunked path computes per-channel log-decay prefixes, causal query/key terms, the beta-folded triangular transform, WY pseudo-keys and pseudo-values, and inter-chunk state carry. It handles a ragged final chunk without a separate tail implementation. Parallelize independent (batch, head) work across the ExecuTorch threadpool. Each worker receives a disjoint slice of one temporary scratch arena, avoiding shared mutable buffers while amortizing allocation across chunks. Reviewed By: billmguo Differential Revision: D112597348
…ization (pytorch#21021) Summary: Reorder the chunked prefill inner loops (steps 1, 4, 5, 6) so the innermost loop runs contiguously over the head dimension (k or v) instead of striding down a column of the state / pv. This lets the compiler autovectorize the now-unit-stride AXPYs; hand-written at::vec was tried and was slower than the compiler output, so the loops stay scalar. Reviewed By: billmguo Differential Revision: D112598714
d103906 to
f9c7ada
Compare
Summary:
Reorder the chunked prefill inner loops (steps 1, 4, 5, 6) so the innermost loop runs contiguously over the head dimension (k or v) instead of striding down a column of the state / pv. This lets the compiler autovectorize the now-unit-stride AXPYs; hand-written at::vec was tried and was slower than the compiler output, so the loops stay scalar.
Reviewed By: billmguo
Differential Revision: D112598714